‹ BackNewsvisual reasoning

visual reasoning

Kaiming He
2026-10-09 07:58:55

Kaiming He’s team introduces VISTA, a visual harness that lets multimodal models revisit past observations

A team led by MIT professor Kaiming He has introduced VISTA, short for "A Visual Harness for Reasoning in an Interactive World," a framework designed to help existing multimodal models preserve and reuse visual experience during long interactive tasks. Instead of converting an environment into text or code and relying on summaries, VISTA stores raw visual frames outside the context window and lets the model pull them back when needed for inspection, comparison, zooming, and pixel-level checks. The paper reports strong results on ARC-AGI-3, an interactive visual reasoning benchmark where rules and goals are not given in advance. With VISTA, Claude Opus 5.0 completed all 25 public games, achieved a relative human action-efficiency score of 100, and used 57.4% fewer game actions than the first-play human baseline. GPT-5.6 Sol also cleared all 25 games and scored 99. The framework is built around three parts: visual observation, lossless visual memory, and active visual inspection. It also uses two text notes, GUIDE.md and WORKING.md, to preserve reusable rules and current task progress across context windows. Beyond ARC-AGI-3, the team tested VISTA on GameWorld, AI GameStore, and BabyVision, reporting gains in browser games, mazes, and connection puzzles. The authors say future validation should move toward embodied tasks in environments that more closely resemble the physical world.

10
Kaiming He’s team introduces VISTA, a visual harness that lets multimodal models revisit past observations
Visual reasoning AI startup Elorian raises $55 million seed round at a $300 million post-money valuation